Poisson-Markov Mixture Model and Parallel Algorithm for Binning Massive and Heterogenous DNA Sequencing Reads

نویسندگان

Lu Wang

Dongxiao Zhu

Yan Li

Ming Dong

چکیده

A major computational challenge in analyzing metagenomics sequencing reads is to identify unknown sources of massive and heterogeneous short DNA reads. A promising approach is to efficiently and sufficiently extract and exploit sequence features, i.e., k-mers, to bin the reads according to their sources. Shorter k-mers may capture base composition information while longer k-mers may represent reads abundance information. We present a novel Poisson-Markov mixture Model (PMM) to systematically integrate the information in both long and short k-mers and develop a parallel algorithm for improving both reads binning performance and running time. We compare the performance and running time of our PMM approach with selected competing approaches using simulated data sets, and we also demonstrate the utility of our PMM approach using a time course metagenomics data set. The probabilistic modeling framework is sufficiently flexible and general to solve a wide range of supervised and unsupervised learning problems in metagenomics.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Title MetaCluster 4 . 0 : A novel binning algorithm for NGS reads andhuge number of species

Next-generation sequencing (NGS) technologies allow the sequencing of microbial communities directly from the environment without prior culturing. The output of environmental DNA sequencing consists of many reads from genomes of different unknown species, making the clustering together reads from the same (or similar) species (also known as binning) a crucial step. The difficulties of the binni...

متن کامل

Title MetaCluster 4 . 0 : A novel binning algorithm for NGS reads

متن کامل

Title MetaCluster 4 . 0 : A novel binning

متن کامل

MetaCluster 4.0: A Novel Binning Algorithm for NGS Reads and Huge Number of Species

متن کامل

Probabilistic insertion, deletion and substitution error correction using Markov inference in next generation sequencing reads

Error correction of noisy reads obtained from high-throughput DNA sequencers is an important problem since read quality significantly affects downstream analyses such as detection of genetic variation and the complexity and success of sequence assembly. Most of the current error correction algorithms are only capable of recovering substitution errors. In this work, Pindel, an algorithm that sim...

متن کامل

ذخیره در منابع من

ذخیره در منابع من قبلا به منابع من ذحیره شده

{@ msg_add @}

با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره شماره

صفحات -

تاریخ انتشار 2016

Poisson-Markov Mixture Model and Parallel Algorithm for Binning Massive and Heterogenous DNA Sequencing Reads

نویسندگان

چکیده

منابع مشابه

Title MetaCluster 4 . 0 : A novel binning algorithm for NGS reads andhuge number of species

Title MetaCluster 4 . 0 : A novel binning algorithm for NGS reads

Title MetaCluster 4 . 0 : A novel binning

MetaCluster 4.0: A Novel Binning Algorithm for NGS Reads and Huge Number of Species

Probabilistic insertion, deletion and substitution error correction using Markov inference in next generation sequencing reads

عنوان ژورنال:

اشتراک گذاری